CMAJ Open
● CMA Impact Inc.
Preprints posted in the last 30 days, ranked by how well they match CMAJ Open's content profile, based on 12 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
McHenry, R. D.; Moultrie, C. E.
Show abstract
Objectives Emergency Department (ED) crowding is an international concern, predominantly caused by 'exit block', the lack of availability of inpatient beds for those requiring admission. The implementation of Flow Navigation Centre Plus (FNC+) services in Scotland aimed to reduce self-presentation to EDs and reduce crowding by providing remote clinical assessment for patients contacting urgent care by telephone and professional-to-professional advice on patient pathways, but their effectiveness is unknown. This study aimed to estimate the effect of board-wide implementation of FNC+ on ED attendances and long waits during the first year of FNC+ operation. Methods Controlled interrupted time series using weekly, publicly reported Public Health Scotland data. The intervention was implementation of the FNC+ in NHS Lanarkshire on 1 April 2024. Counts were summed across constituent sites and percentages derived from board totals. Co-primary outcomes were ED attendance volume and the proportions of attendances spending more than 4, 8 and 12 hours in the department. Segmented regression was fitted with contemporaneous control boards, seasonal terms, and accounted for autoregression. Results 118 pre-intervention and 52 post-intervention weeks were analysed across all 3 EDs in the implementing board. Attendances showed no detectable step change (+1.20%; 95%CIs -0.66 to +3.10) relative to the counterfactual. The estimated effect increased across follow-up, however, changing by +3.95% over 52 weeks (95% CI +0.36 to +7.67%). There was no significant step change in the proportion of attendances waiting more than 4 hours following the intervention (+1.74%; 95%CIs -0.71 to 4.20%). Some transition and structural sensitivity analyses demonstrated significant deteriorations in ED performance, and increased attendances, in the year following implementation, and none demonstrated improvements. Conclusions Board-wide implementation of a Flow Navigation Centre Plus was not associated with a step change in ED attendances or in long waits, but there is some evidence that attendances increased and long waits increased in the year following implementation. Their provision of supply-sensitive care is a possible mechanism. Additionally, given their action at the point of input, aiming to divert patients from ED attendance, it is unlikely that such services could relieve a constraint due to exit block, the availability of inpatient care for those requiring admission.
Parpia, A.; Wright, J.; Gharouni, A.; Thampi, N.; Fitzpatrick, T.
Show abstract
Background: Respiratory syncytial virus (RSV) remains a leading cause of hospitalization in infancy, with severe outcomes influenced by both contact patterns and passive immunity. Non-pharmaceutical interventions (NPIs) during the COVID-19 pandemic suppressed RSV circulation and reduced opportunities for maternal immune boosting, potentially altering protection among newborns. We evaluated whether incorporating time-varying maternal immunity improves the ability of an age-structured transmission model to predict post-pandemic RSV hospitalization patterns in infants. Methods: We analyzed population-based RSV hospitalizations among Ontario (Canada) infants (<1 year) from July 2, 2017 to June 25, 2024, using linked administrative databases. A deterministic compartmental model across seven age classes was calibrated against pre-pandemic data using Latin Hypercube Sampling. We compared a model incorporating time-varying contact rates alone against a specification that additionally included time-varying maternal immunity. Results: Both specifications accurately reproduced pre-pandemic seasonality and macro-level post-pandemic resurgence features. The constant maternal immunity model showed slightly better accuracy in capturing the 2021/22 peak compared to the time-varying maternal immunity specification. However, both qualitatively captured the continued near-absence of RSV and the observed peak was captured within the 95% credible intervals. While both models precisely captured the timing and overwhelming surge of admissions that occurred in 2022/23, they failed to capture the premature peak timing and magnitude in 2023/24. Conclusions: Incorporating time-varying maternal immunity did not improve model accuracy post-pandemic. While maternal protection is essential for evaluating infant immunizations, population-level contact shifts primarily shaped post-pandemic RSV seasonality, indicating that models must account for these mechanisms of RSV transmission dynamics.
Rees, N.; Angouri, J.; Ting, S. S. P.; Booker, M.; Nadeem, L.; Williams, L.; Rawlinson, D.
Show abstract
Background Decision-making in emergency services involving the allocation of scarce resources is a key challenge for large, complex organisations required to prioritise demand against multiple, often competing criteria. Emergency Medical Dispatch is a case in point, where Enhanced and Critical Care Teams (ECCTs) represent a scarce and lifesaving clinical resource. Despite its operational and system-level significance, the allocation of ECCTs remains under-researched. Methods We conducted a methodological development study using the P.A.T.H.S. framework (Participants, Artefacts, Transition Stages, Historicity, Setting). We designed and piloted this in our previous work under the 999 R.E.S.P.O.N.D. project, which examined the decision-making process for ECCT dispatch. We applied P.A.T.H.S. to 17 dispatch cases (comprising 100 decision-making episodes). We analysed five data sources: recordings of emergency calls and internal dispatch-related interactions, sequence-of-events records, policy documents, and ethnographic observations. A four-phase analysis--indexing & data mapping, transcription & coding, charting, and synthesising & outputs--was undertaken taking Interactional Sociolinguistics as the theoretical approach and methodology. Results P.A.T.H.S. enables the mapping of non-linear, multifactorial textual trajectories across human and non-human actors. The case example presented herein illustrates how information on key risk indicators (e.g. mechanism of injury) were often delayed, fragmented, or lost between the caller, call-handler, and written records. P.A.T.H.S. provides a framework and analytical tool capturing the textual trajectory of information flow, and trace how dispatch decision making unfolds. We subsequently developed a template and codebook for other researchers to further study complex decision making in multi-actoral systems using a textual trajectory approach. Conclusion This methodological development work demonstrates the potential of P.A.T.H.S. to capture and clarify complex decision-making processes. P.A.T.H.S. offers a practical and theoretically grounded framework for future research, training, and policy that addresses risk points in communication between oral and written forms among teams of actors, to support optimal deployment of scarce resources.
Hauser, K. A.; Degesys, N. F.; Isaacs, E. D.; Tang, M.; Swartzberg, J.; Panopulos, V.; Martin, A. M.; Liu, V. X.; Schlessinger, D.; Samady, N. A.; Malhotra, R.; Plimier, C.; Hadadianpour, A.; Erickson, M. D.; James, T.; Rogers, S.; Adler-Milstein, J.; Thombley, R.; Rosenthal, S.; Harris, A. R.; Hardy, J.; Raven, M.; Singh, M.; Kim, C.; Perry, R.; Clevenger, E.; Carvajal, C.; Babino, D.; Gray, A.; Shapiro, M.; Chan, T.; Allore, H.; Meeker, D.; Tomasino, D.; Grogan, E. F.; Pepper, A.; Wellons, M.; Hwang, U.
Show abstract
Background: Three San Francisco health system emergency departments have developed Geriatric Emergency Department (GED) models of care programs supporting and providing care for emergency department (ED) patients at risk for or living with dementia. Each system recognized: 1) the high proportion of older adult ED patients and those at risk for dementia, 2) the need to identify cognitive impairment in older adult ED patients, 3) the importance of developing approaches to connect older adult ED patients and their care partners with resources and diagnostic specialty services. Methods: We describe how each hospital adopted and implemented pragmatic GED models of care to support and improve care for ED patients at risk or living with dementia. We also report the proportion of ED encounters made by patients with dementia histories and the number of these reached by GED programs. Results: Three San Francisco hospitals (a tertiary care, critical access, and large integrated health system-community ED) independently implemented GED programs to support and enhance emergency care for patients living with dementia. Each uses screening and assessment tools to identify patients at risk for cognitive impairment. Each captures screening and assessment data to facilitate care and resources for post-discharge care, ensuring coordinated transitions and support for older adults. Programs varied by target patient population age and staff and resource allocation to support program goals. Site-specific pathways differed by location, patient populations, and support from geriatrics, emergency medicine, palliative medicine, neurology, psychiatry, pharmacy, referral processes, and/or pastoral care. Conclusions: Developing GED care interventions that facilitate care for patients at risk of or living with dementia is possible and sustainable when the pathway aligns with health system leadership goals through persistent value demonstration, communication, and promotion. Ultimately, developing and disseminating models of GED care is designed to address geriatric syndromes inclusive of dementia care through continuous quality improvement.
Henry, K.; Smith, B. A.; Holden, D. N.; Smith, S. E.; Heavner, M. S.; Chen, Z.; Chen, X.; Devlin, J. W.; Murphy, D. J.; Martin, G. S.; Burden, M.; Murray, B.; Sikora, A.
Show abstract
Background: While critical care pharmacists (CCPs) are broadly associated with improvements in outcomes for critically ill patients, operationalizing staffing in the intensive care unit (ICU) requires further study. The purpose of this evaluation was to determine the relationship of a CCP on interprofessional rounds for weekday admissions of ICU patients on patient-centered outcomes. Methods: This post-hoc analysis of the Optimizing Pharmacist-Team Integration for ICU Patient Management (OPTIM) study included adults admitted to an ICU on a weekday in the multicenter observational study. The primary outcome was in-hospital mortality. The primary exposure was level of comprehensive medication management (CMM) during the first 24 hours of ICU stay. A secondary exposure was pharmacist-to-patient ratio. Multivariable generalized estimating equations (GEE) were used to estimate associations between mortality and patient, ICU, and institution variables. Fine-Gray sub-distribution hazards regression estimated hazard of discharge alive (HDA) from the ICU and hospital and hazard of extubation alive. Results: 21,835 patients met inclusion criteria, and 76.1% of patients had CMM delivered on interprofessional rounds. Patients who had no CMM on the first ICU day had an increased risk of mortality of 23% (Odds Ratio (OR) 1.23, 95% Confidence Interval (CI) 1.04-1.46, p=0.02) compared to those who received CMM on interprofessional rounds. Patients with no CMM also had decreased HDA from the ICU and hospital and decreased hazard of extubation alive. No difference was seen in any outcomes when comparing other levels of CMM (CMM delivered outside of interprofessional rounds or abbreviated CMM) compared to CMM delivered on rounds. Conclusions: Absence of pharmacist CMM on the first day of ICU stay for patients with weekday admission was associated with an increased risk of in-hospital mortality, but no difference was seen in other levels of CMM: this signal supports further investigation in prospective analysis.
Plagenz, J.; Lin, A.; Harlow, T.
Show abstract
Background: Timely carbidopa-levodopa administration is a recognized inpatient safety priority in Parkinson disease, and mistiming is common, but where in the medication-use process it arises is uncharacterized. Objectives: To localize where inpatient mistiming arises and where to target intervention. Methods: In a single-center retrospective analysis of hospitalized adults with Parkinson disease on home carbidopa-levodopa, each dose's administration time was compared with the individualized home schedule. Mistiming was defined a priori as more than 15 minutes from the home time (Parkinson's Foundation Hospital Care Standard 2). We characterized the deviation distribution, tested whether administrations tracked the schedule or the standard grid, and examined length-of-stay and readmission. Results: Across 947 doses in 101 patients, ordering was accurate, yet 62.9% (596 of 947) missed the home time by more than 15 minutes and 99% of patients had at least one mistimed dose. Administrations tracked the individualized schedule almost exactly (Pearson r 0.98), not the standard grid: only 10% fell within 15 minutes of the default times, and the median dose sat 24 minutes from its home time but 76 from the nearest default. Deviation was symmetric drift (median absolute deviation 24 minutes; 16.5% beyond 60 minutes). Conclusions: Mistiming in this study reflected imprecise bedside execution, not ordering or a mismatch between fixed rounds and individualized regimens. These findings may point medication-safety efforts toward protecting bedside administration as complementary redesigning orders.
McHenry, R. D.; Saunders, A.; Ahmad, F.; Mackay, D.
Show abstract
Background Emergency Department (ED) crowding is an international crisis primarily driven by exit block. Point of care (POC) cardiac biomarker testing and reduced sampling intervals have been proposed to mitigate crowding by improving throughput, but whole-ED operational impacts remain poorly understood, and evaluations often rely on vulnerable observational designs. This study aimed to assess whether introducing POC high-sensitivity troponin testing and reduced sampling intervals changed whole-ED flow metrics, and to test the robustness of interrupted time series (ITS) methodology in this setting. Methods A multi-centre controlled interrupted time series (CITS) across two large urban intervention EDs and one untreated control ED in Glasgow, UK. The intervention combined whole-blood POC high-sensitivity troponin testing with a reduction in sampling intervals from 3 to 2 hours. Outcomes included daily ED admissions, mean occupancy, maximum occupancy, and mean length of stay. Analyses used a window of 120 days either side of each implementation date. Effects were evaluated using segmented ITS models, with and without controls, with permutation tests against 147 pre-intervention placebo dates. The minimum detectable effects of a similar study, applied to a national dataset, were simulated. Results Across 483,412 presentations to the intervention sites, the intervention produced no statistically significant change in any whole-ED flow metric against the untreated control at either site. Analysed alone, one intervention site appeared to show reductions in mean occupancy (-6.08, 95% CI -12.04 to -0.12) and maximum occupancy (-7.60, -14.47 to -0.73); the untreated control department produced reductions in the same direction at the same date, and both estimates attenuated to the null once the control was applied. Under a pre-specified 14-day transition specification the reductions in the untreated department reached statistical significance while those at the treated site did not. The study was limited by power due to the study window and limited control pool. Simulation demonstrated that a national dataset has the potential to provide operationally feasible and clinically important findings. Conclusion POC cardiac biomarker testing and reduced sampling intervals did not detectably improve whole-ED flow, though the design was underpowered. More importantly, uncontrolled ITS designs are highly vulnerable to confounding in complex healthcare systems; evaluations of operational interventions must utilise concurrent controls, and routinely report falsification tests.
Chaudhry, R.; Chen, Z. S.
Show abstract
Background: Mortality prediction models often combine early electronic health record data, but the relative prognostic value of baseline vulnerability, physiological severity, treatment exposure, and procedure burden remains unclear. Objective: To compare routinely available first-24-hour clinical domains for visit-level 30-day mortality prediction and assess whether domain-level patterns replicated in MIMIC-IV. Methods: We used CHoRUS, an OMOP-formatted acute-care dataset, with independent domain-level replication in MIMIC-IV. CHoRUS included 22,098 visits among 5,892 unique patients, with 1,004 30-day mortality events and 4.5% mortality prevalence. MIMIC-IV included 23,000 acute-care visits among 10,006 unique patients, with 819 events and 3.6% prevalence. Across both datasets, 45,098 visits and 15,898 unique patients were analyzed. Predictors were restricted to the first 24 hours after visit start. Performance was evaluated using AUPRC, AUROC, Brier score, calibration, sensitivity at 90% specificity, highest-risk 10% analyses, decision-curve analysis, and SHAP summaries. Because 30-day mortality was infrequent, the classification task was class-imbalanced. Accordingly, AUPRC was interpreted relative to the prevalence-based no-skill baseline, rather than as an absolute measure alone. Results and Conclusion: Physiological severity produced the largest improvement beyond baseline in CHoRUS, with median AUPRC 0.38 and median AUROC 0.86, and showed the same primary domain-level pattern in MIMIC-IV. Treatment exposure and procedure burden provided smaller gains. In CHoRUS, the best pairwise model combined baseline, physiological severity, and procedure burden features, with median AUPRC 0.41; the all-domain model was slightly lower, with median AUPRC 0.40 and median AUROC 0.86. In MIMIC-IV, the all-domain model had the highest median AUPRC, 0.25, only modestly above the best pairwise model. First-24-hour physiological severity features therefore provided the most consistent prognostic information across datasets, supporting parsimonious, clinically interpretable acute-care risk models centered on high-quality early physiological data.
Sugawara, H.
Show abstract
Background: Whether corrective actions documented in medical safety incident reports rely on individual vigilance ("Safety-I") or on structural, system-level intervention ("Safety-II") has not been quantitatively evaluated on a national scale in Japan. We developed an automated classification pipeline to assign corrective-action free-text to a 7-level maturity scale (L0-L6) and computed two summary indices: the Safety Measure Quality Profile (SMQP), the full L0-L6 distribution, and the System-based Safety Measure Rate (SSMR), the proportion of non-L0 records classified L3-L6. Methods: We analyzed all 11,507 corrective-action free-text entries from the 2010 release of Japan's national medical accident and near-miss reporting database (Japan Council for Quality Health Care, JCQHC), comprising 8,804 near-miss (Hiyari-Hatto) and 2,703 accident (Jiko) reports. Records were classified using a five-stage hybrid pipeline: an expert-developed rule dictionary, TF-IDF + k-nearest-neighbor matching, cosine-similarity matching, a two-tier large-language-model (LLM) classifier, and a conservative priority-cascade fallback. SSMR was compared between near-miss and accident reports using a chi-square test, Wilson 95% confidence intervals, Cramer's V, and the risk difference (RD), against pre-specified minimal clinically important difference (MCID) criteria of RD >= 2 percentage points and Cramer's V >= 0.10. Results: Every record received a definitive L0-L6 label (0% unresolved). Overall, 16.6% of records were unclassifiable (L0); among the 9,599 classifiable (non-L0) records, individual-vigilance actions (L1) predominated (54.6% of all records), and only 11.82% (95% CI, 11.19-12.49%) met the SSMR criterion (L3-L6). SSMR was higher for accident reports than for near-miss reports (18.12% [95% CI, 16.69-19.64%] vs. 9.[95% CI,46% [95% CI, 8.80-10.17%]; RD = 8.66 percentage points; Cramer's V = 0.120; chi-square(1) = 136.97001), exceeding both pre-specified MCID thresholds. Conclusions: In this interim single-year analysis, the large majority of documented corrective actions in Japanese medical safety reports remained individual-vigilance-based rather than system-based, with accident reports showing a substantively, rather than merely statistically, higher proportion of system-based actions than near-miss reports. These findings support the feasibility of large-scale automated assessment of corrective-action quality and provide the rationale for the planned 16-year longitudinal analysis.
McHenry, R. D.; Caesar, D.; Clarke, B.; Mackay, D.; Pell, J.
Show abstract
Objectives Emergency department (ED) crowding is recognised as an important public health concern internationally, and is driven principally by exit block, the shortage of inpatient beds for patients requiring admission. This study aimed to evaluate whether a complex intervention targeting hospital occupancy improved ED patient flow, and quantified the change in attendances. Methods A controlled interrupted time series using weekly, publicly reported Public Health Scotland data from 1 January 2022 to 1 February 2026. The multi-component intervention focused on reducing hospital occupancy and included additional adult social care funding; engagement with regional social care providers; accelerated implementation of the Discharge without Delay programme; re-evaluation of whole-hospital escalation thresholds and response; resource and data supporting inpatient department reductions in length of stay; and additional investment in remote clinical assessment. The intervention commenced at a large tertiary ED on 01 February 2025. Primary outcomes were the proportions of attendances spending [≥]4, [≥]8 and [≥]12 hours in the ED. The secondary outcome was attendance volume. Segmented regression was fitted with a contemporaneous control series, seasonal terms and autoregressive moving average errors. Long waits were additionally illustrated as potentially avoided deaths. Results The analysis covered 161 pre-intervention and 52 post-intervention weeks. Relative to pre-intervention levels, the proportion of attendances waiting over 4 hours fell by 10.4% (95% CI 1.6 to 19.2%), by 16.4% (95%CI 1.3 to 31.5%) over 8 hours and by 24.3% (95%CI 2.6 to 46.1%) over 12 hours. Using established associations between long ED waits and excess mortality, by one-year the intervention was potentially associated with 54 fewer excess deaths (95%CI 19 to 93). Attendances rose by 3.8% (95%CI 1.3 to 6.4%) against the counterfactual. Conclusions A complex intervention targeting hospital occupancy was associated with a reduction in long ED waits despite rising attendances. Interventions addressing hospital occupancy can meaningfully improve ED crowding.
Presanis, A. M.; Nyberg, T.; Rolfes, M. A.; Quinot, C.; Goudie, R.; Whitaker, H. J.; Elson, W. H.; Byford, R.; Mikdashi, T.; Wong, J. Y.; Andrews, N.; Villar, S. S.; Cowling, B. J.; Charlett, A.; Dabrera, G.; Pebody, R.; Lopez Bernal, J.; de Lusignan, S.; De Angelis, D.
Show abstract
Influenza surveillance has typically been carried out using influenza-like illness (ILI) rates and proportions of laboratory tests positive for influenza as metrics to monitor, with sample sizes for the number of tests to carry out based on the precision of the resulting estimate of proportions positive. The transition out of the Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) pandemic period has encouraged the establishment of integrated surveillance of respiratory pathogens, in the context of multiple surveillance objectives, as set out by WHO in its revised integrated surveillance guidance and Mosaic Respiratory Surveillance Framework. These objectives include outbreak detection, situational awareness and intensity evaluation, among others. We illustrate how to design respiratory surveillance in primary care, by considering multiple surveillance objectives for different metrics of different types of respiratory pathogen circulation seasons in England, the USA and Hong Kong. We focus on a proxy of influenza activity as a metric to compare between these countries/regions. Taking advantage of England's integrated sentinel primary care surveillance system, we propose further metrics to monitor: a proxy of respiratory activity, novelly defined as the product of an acute respiratory infection (ARI) consultation rate and the proportion of tests positive for \emph{at least one pathogen}; pathogen-specific ARI-based activity proxies for more detailed monitoring of influenza and SARS-CoV-2; and integrated monitoring of proportions positive for all pathogens tested. We use a simulation approach to determine sample sizes by optimising either the probability of, or time to, detection of different events in monitored metrics, according to the different surveillance objectives. We find that sample sizes to maximise detection probabilities or minimise detection times vary by metric, objective, event and country/region. At a national level, the current sample sizes used are sufficient to detect most events in most weeks for both the USA and Hong Kong, but for England the numbers of swabs taken for ILI consultations may not be sufficient in all weeks, particularly at the start of the season when outbreak detection is important. However, broadening the criteria for swabbing to acute respiratory symptoms does allow for sufficient sample sizes.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Kim, S. S.; Zissette, S. Z.; Van Meter, C.; Shiiba, M.; Bruck, M.; Tippett, A.; Kamidani, S.; Benkeser, D.; McQuade, E. R.
Show abstract
Importance: Maternal vaccination and long-acting monoclonal antibodies are now available in the U.S. to prevent RSV. Long-acting monoclonal antibody administration in the U.S. commonly occurs after hospital discharge in outpatient settings, leaving some infants unprotected early in life when severe RSV risk is highest. Comparative effectiveness between the two interventions and whether delays affect effectiveness estimates have not been quantified. Objective: To evaluate the effectiveness of infant long-acting monoclonal antibody strategies and a maternal vaccination strategy, each compared to no intervention, and the comparative effectiveness of intervention strategies when accounting for real-world delays in monoclonal antibody receipt. Design: Cohort study using target trial emulation to compare four strategies for prevention of RSV-related outcomes. Setting: The U.S. between 2023 and 2025 using a nationwide database of employer-sponsored commercial insurance claims. Participants: 120,586 commercially insured mother-infants, whose infants were born in the U.S. during the 2023-2024 or 2024-2025 RSV season. Infants who could not be paired with their mother's record, did not enroll in commercial insurance within 75 days from birth, received palivizumab, and had an implausible birth date were excluded. Interventions: Comparison of four RSV prevention strategies: (i) maternal RSVpreF; (ii) long-acting monoclonal antibody given within the first week of life (mAb as intended); (iii) long-acting monoclonal antibody given within a six-month grace period from birth (mAb within grace period); and (iv) a control. Main outcomes and measures: Effectiveness against first RSV-associated hospitalization and medically-attended RSV illness was summarized using adjusted hazard ratios (aHR) and estimated using an inverse propensity weighting approach, with weights accounting for maternal age, maternal comorbidities affecting pregnancy, obstetric and newborn complications, season, region, and birth timing relative to October 1. A weighted Kaplan Meier estimator was used to estimate strategy-specific cumulative incidence of RSV outcomes over time. Results: In the first five weeks of life, the mAb within grace period strategy doubled the hazard of RSV hospitalization (aHR: 2.0 [95% CI: 1.0-4.9]) and increased the hazard of medically-attended RSV (aHR: 1.6 [95% CI: 1.0-2.7]) compared to the maternal RSVpreF strategy. The hazard for RSV hospitalization was similar for the mAb as intended strategy compared to the maternal RSVpreF strategy (aHR = 0.9 [95% CI: 0.3-1.9]). Conclusions and relevance: RSVpreF and monoclonal antibodies were similarly effective when monoclonal antibodies were administered close to birth, but when accounting for real-world delays in monoclonal antibody receipt, the maternal RSVpreF strategy was more effective than the mAb within grace period strategy.
Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;
Show abstract
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.
Gao, Q.; Hayhoe, B.; Cicek, M.; Greenfield, G.; Otis, M.; Misirli, G.; Luisa Neves, A.; Majeed, A.; Aylin, P.; Bottle, A.
Show abstract
Objectives To assess the concurrent and lagged associations between quality of primary care and planned and unplanned secondary care use for patients with multimorbidity, examining the modifying role of frailty. Design A retrospective cohort study Setting This population-level analysis included 468,172 patients with multimorbidity in England from the Discover research platform (April 2022-March 2024). Participants Patients with multimorbidity Main outcome measures We used principal component analysis to combine a set of quality indicators (QIs) and assessed the impacts of QIs on both planned and unplanned care. Results Generally, patients with higher QI attainment also had higher likelihood of planned (outpatient visits) and unplanned care (emergency admissions and ED visits) utilisation. There was a lower lagged odds of elective hospital admissions in the following 12 months among those with higher attainment of multimorbidity-specific QIs (OR=0.94, 95%CI 0.93-0.95). In the complex multimorbidity cohort ([≥]3 conditions), multimorbidity-specific QIs were longitudinally associated with lower odds of elective admissions (OR=0.94, 95%CI 0.92-0.95) and outpatient visits (OR=0.96, 95%CI 0.95-0.98), while generic QIs were related to lower odds of outpatient non-attendance (OR=0.95, 95%CI 0.91-0.99). In non-frail patients with multimorbidity, multimorbidity-specific QIs were longitudinally associated with reduced odds of outpatient visits (OR=0.98, 95%CI 0.97-0.99), elective admissions (OR=0.92, 95%CI 0.90-0.94) and prolonged elective hospital stay (IRR=0.94, 95%CI 0.89-0.99). Conclusions Attainment of generic and multimorbidity QIs was generally associated with slightly increased planned and unplanned care. However, patients for whom we identified higher attainment of multimorbidity-specific QIs had lower odds of elective admissions and outpatient visits, especially for those with complex multimorbidity. Our research suggests that the quality of primary care may influence patients' use of secondary care, with the potential to improve care for people with multimorbidity and warrant further investigation into management strategies.
Jaganath, D.; Ilavarasan, V.; Wong, R.; Chitnis, A.; Murrill, M. T.
Show abstract
Context: Most individuals in the United States have commercial health insurance, yet costs for tuberculosis (TB) care have focused on the public sector. Objective: To quantify 12 month all cause healthcare costs and identify predictors of expenditure among commercially insured persons with TB disease in the United States. Design/Setting: Retrospective cohort study using Merative (TM) MarketScan (R) Commercial Claims Database (2013 to 2018). Participants: Adults 18 years old with TB disease Main Outcome Measure: Total 12 month all cause healthcare costs (outpatient, inpatient, pharmacy) were calculated from the date of diagnosis. Adjusted cost ratios (aCR) were estimated using a Gamma generalized linear model. Results: We included 303 individuals diagnosed with TB disease, median age 46 years, 158 (52%) male, 16 (5%) with HIV, 12 (4%) with hepatitis B (HBV), and 13 (4%) with a drug use disorder. Mean total 12-month costs were $32,404 (median $8,075; SD $78,829). Median 12-month costs were substantially higher among persons with any comorbidity (HIV, HBV, hepatitis C (HCV), alcohol use disorder, drug use disorder, or Charlson score >0) compared to those without ($11,930 [IQR $4,194 to $36,073] vs $3,385 [IQR $1,506 to $8,609]; p<0.001). HIV coinfection and drug use disorder were the strongest independent predictors. HIV coinfection was associated with 4.7 fold higher costs (aCR 4.70, p<.001), driven predominantly by pharmacy expenditure (aCR 16.4). Drug use disorder was associated with 3.2 fold higher costs (aCR 2.62, p=.03). Comorbidity burden was a continuous independent predictor (aCR 1.36 per Charlson point, p<.001). Conclusions: Healthcare costs are high among persons with TB who have commercial insurance, and are further increased with comorbidities including HIV coinfection and drug use disorder. Improved screening, care coordination and management of TB and high risk comorbidities could yield significant cost savings.
Smith, S. J.; Lemoine, D.
Show abstract
Objective: To assess the efficacy of an executive peer coaching program, Charting Champions Program (CCP), in helping physicians manage their administrative workload, thereby improving time management, workflow and well-being. Findings: In this longitudinal survey study, physicians self-reported significant improvements in completing charting and administrative paperwork during their clinical day. Physicians reported significant improvements in mental, cognitive and emotional states after the program. Meaning: The Charting Champions Program is an effective intervention that supports physicians in problem-solving the administrative burden of their clinical day, improving workflow efficiency, completing administrative requirements during clinical hours, and enhancing work-life balance and personal satisfaction. Background: Physicians are subject to high levels of mental, physical, and emotional stress, partly due to increasing administrative burdens. Online coaching is a proven intervention to help physicians improve workflow efficiency, reduce administrative burden and improve job satisfaction. Design: This voluntary longitudinal survey took place between 2020 and 2023. Physicians were asked to complete a survey at program entry and again 30-90 days after program completion. The survey consisted of 14 Likert scale questions, and a final sample of 280 physicians completed both surveys. Intervention: CCP contains modules that teach workflow improvements for clinical days, including timely charting, administrative task workflow, managing patient consultations and reducing interruptions. Interventions include self-paced modules, live coaching, recordings and an online peer community. Results: Post-CCP physicians reported a significant decrease in hours spent charting (P<0.0001) and completing clinical paperwork outside of clinical hours (P<0.006). Physicians also reported a decrease in work-related dread (P<0.001), feelings of burnout (P<0.001), and thoughts of quitting due to administrative burdens (P<0.001). Physicians felt more focused at work (P<0.001), felt more in control of the clinical day (P<0.001), and rated their mental energy at work higher (P<0.001). The program did not affect the number of patients seen in a full clinical day (P > 0.918). Conclusion and Relevance: The CCP reduces the time physicians spend on tasks outside of clinical hours, increasing free time without decreasing the number of patients seen per day.
Roberts, L.
Show abstract
Objective. Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify to optimal approach. Methods. Twenty referral scenarios spanning the urgency spectrum, based on real referrals were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded, to produce a consensus reference standard. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per condition). The experiment was run with a simple prompt and repeated with a advanced prompt supplying explicit triage expectations and worked examples. Results. All 2760 attempts returned valid categories. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho=0.42; P=.047) and accuracy tracked cost. Advanced prompting minimised between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P=.01), abolished the size-accuracy association (rho=-0.05; P=.83) and removed the accuracy-cost relationship. Leading models matched expert consensus on most cases, within or above the range reported for human triage. Under-triage errors persisted with some LLMs. Conclusion. Contemporary LLMs categorise rheumatology referral urgency as well or better than published human triage systems. Advanced LLM prompting methods substitute for the reasoning capability of larger models, suggesting that LLM performance on this task may not require the most expensive models. The tools to automate this administrative task appear to already exist. Strong candidate LLMs that might serve a production ready solution have been identified.
Aldis, R.; Wang, S.; Sage, M.; Metzmaker, M.; Galvin, H.
Show abstract
Ambient artificial intelligence scribes are being increasingly used in healthcare to improve efficiency and reduce provider clinical documentation burden, yet their performance across linguistically diverse patient populations is not well characterized. We conducted a retrospective analysis of 54,160 outpatient encounters within a U.S. safety net health system to evaluate the performance of an artificial intelligence documentation tool in English and non-English clinical encounters, and in encounters where an interpreter or bilingual provider was present. Documentation performance was measured by the percentage of words in the final note that were generated by the ambient AI documentation tool and not edited by the provider. Associations between language factors and documentation performance were measured using Generalized Estimating Equations with exchangeable correlation structures to account for clustering of multiple encounters within unique patients. Univariable models were fitted to estimate the odds of adequate performance by language and interpreter modality, and a multivariable interaction model was used to evaluate within-language differences between bilingual providers and interpreter-mediated encounters. Non-English encounters were 21% to 25% less likely than English encounters to achieve the same performance threshold. There was no significant difference in generative documentation performance between interpreter-mediated and bilingual provider encounters. These findings underscore the importance of equity-focused evaluation and multilingual model refinement to ensure that artificial intelligence documentation benefits are distributed fairly across diverse patient populations.